Back

BioData Mining

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match BioData Mining's content profile, based on 22 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
AI as a signal assessor - Can a Large Language Model perform causality assessment on a case series?

Shenoy, A.; Zekarias, A.; Viklund, A.; Mitchell, J.; Barrett, J.; Sandberg, L.; Meldau, E.-L.; Taavola-Gustafsson, H.

2026-06-29 pharmacology and therapeutics 10.64898/2026.06.26.26356656 medRxiv
Top 0.1%
10.9%
Show abstract

Background Large Language Models (LLMs) are increasingly explored for pharmacovigilance tasks, including information extraction, case documentation, and single-case causality assessment. However, their ability to support causality assessment at the case series level -- a complex, time-intensive task requiring clinical reasoning across multiple reports -- remains unexplored. Objective To investigate how a large-scale general-purpose LLM can support pharmacovigilance professionals in assessing causality in a case series, and to explore how prompt design influences the quality of the model's reasoning. Methods GPT-4o was used to assess causality for five drug - adverse event combinations, using an adaptation of the Bradford Hill viewpoints for case series assessment. The combinations represented varying drugs and vaccines, adverse events, and case series sizes (5-402 reports). One combination served as a negative control. Structured prompts were iteratively developed and refined using one combination, then applied to all combinations. LLM-generated assessments for each viewpoint were qualitatively evaluated by human annotators for accuracy (precision), and the LLM's coverage of key aspects from the original signal text was assessed for one combination (recall). Results Across all five combinations, annotators agreed with 79-92% of the LLM's output sentences. Full disagreement was consistently low (3-7%), with errors typically involving misinterpretation of complex report details rather than outright fabrication. Prompt design substantially influenced output quality; providing Bradford Hill viewpoint descriptions, including case series data, and adding explicit anti-hallucination instructions improved specificity and grounding. For the recall assessment, 15 of 23 key segments from the original signal text were reflected in the LLM output. The overall summary assessments demonstrated balanced reasoning, correctly distinguishing between positive safety signals and the negative control, and provided a coherent synthesis suitable as a starting point for human assessors. Conclusions LLMs have the potential to generate contextually nuanced and largely accurate preliminary causality assessments of case series aligned with the Bradford Hill viewpoints, with a low but non-zero hallucination rate. These findings support LLMs as a tool to augment, not replace, expert judgment in signal assessment. Future work should address larger and more diverse signal sets, improved evaluation frameworks for generative output, and the integration of pre-computed summary statistics to reduce errors.

2
Precision survival estimation in acute myeloid leukemia using evolutionary learning-derived microRNA signature

Yerukala Sathipati, S.; Agustriawan, D.; Gopireddy, N. S. R.; Popat, A.; Moat, L.; Aimalla, N.; Elugoti, M. R.; Kampa, S. A.; Sharma, P.; Ho, S.-Y.; Sharma, R.

2026-05-26 bioinformatics 10.64898/2026.05.22.727196 medRxiv
Top 0.1%
10.2%
Show abstract

BackgroundAcute myeloid leukemia (AML) remains the most lethal acute leukemia in adults, with 5-year overall survival below 32% despite recent advances including venetoclax-, FLT3-, IDH1/2-, and Menin-targeted therapies. Clinical outcomes remain highly heterogeneous across patients, highlighting the need for robust molecular biomarkers capable of improving prognostic precision. MicroRNAs (miRNAs) are critical regulators of hematopoietic differentiation, apoptosis, and therapeutic resistance and are differentially expressed across AML subtypes. However, their clinical translation has been limited by high dimensionality, feature redundancy, and relatively small cohort sizes. MethodsWe developed and evaluated the AML Survival Estimator (AMLS), an inheritable bi-objective combinatorial genetic algorithm integrated with support vector regression (SVR), using TCGA-LAML miRNA expression profiles (n = 156). AMLS was benchmarked against ten widely used machine-learning approaches, including penalized regression, tree-based ensembles, support-vector regression, k-nearest neighbors, and multilayer perceptron models. Performance was assessed using stratified cross-validation with Pearson correlation (R), Harrells concordance index (C-index), and mean absolute error (MAE). Functional characterization of the derived miRNA signature was performed through consensus target integration followed by pathway enrichment, gene ontology analysis, network reconstruction, and Kaplan-Meier risk stratification. ResultsAMLS achieved superior prognostic performance with pooled out-of-fold metrics of Pearson R = 0.86, C-index = 0.788, and MAE = 7.49 months, substantially outperforming all comparator models. Restricting analyses to the AMLS-derived 28-miRNA signature improved all baseline learners by approximately 2-4-fold, with the multilayer perceptron achieving R = 0.674; however, none matched the native AMLS framework, indicating that the evolutionary optimization strategy contributes predictive information beyond feature selection alone. The prognostic signature included biologically established AML-associated miRNAs, including hsa-miR-191, hsa-miR-29c, hsa-miR-125b, hsa-miR-148a, hsa-miR-15b, hsa-miR-10b, and hsa-miR-30c, linked to DNA methylation, apoptosis, cell-cycle regulation, and oncogenic Wnt/MAPK signaling pathways. Functional analyses demonstrated significant enrichment of canonical AML-associated pathways, including p53, PI3K-AKT, TGF-{beta}, JAK-STAT, FoxO, and hematopoietic lineage signaling. ConclusionsOur findings demonstrate that evolutionary learning integrated with SVR can recover a compact and biologically interpretable miRNA prognostic signature that substantially outperforms conventional machine-learning approaches for AML survival prediction. The identified miRNA network converged on key leukemogenic pathways involved in apoptosis, cell-cycle regulation, and oncogenic signaling, supporting both the biological relevance and prognostic utility of the framework. Given the minimally invasive and quantitatively scalable nature of miRNA profiling, this approach may provide a practical molecular adjunct for improving prognostic assessment and precision medicine strategies in AML. Abstract FigureSchematic overview of the AMLS framework. Left: acute myeloid leukemia, a clonal hematological malignancy with persistent prognostic heterogeneity. Middle: AMLS couples an evolutionary learning-based feature selection algorithms to support vector regression for miRNA-based survival modeling. Right: AMLS recovers a 28-miRNA prognostic signature that predicts overall survival with Pearson R = 0.86 and MAE = 7.5 months. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=86 SRC="FIGDIR/small/727196v1_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@11ead1org.highwire.dtl.DTLVardef@4f5c19org.highwire.dtl.DTLVardef@277de1org.highwire.dtl.DTLVardef@b95c9a_HPS_FORMAT_FIGEXP M_FIG C_FIG

3
Personalizing Suicide Risk Assessment: Machine Learning Extraction of Cross-Modal Interactions Between Psychosocial and Demographic Factors in Veterans

Levis, M. E.; Shiner, B.; Dimambro, M.; Rozema, L.; Ayandeh, S.; Diallo, A. B.; Zhou, Y.; Li, S.; Wu, W.; Gui, J.; Levy, J. J.

2026-06-18 psychiatry and clinical psychology 10.64898/2026.06.16.26355796 medRxiv
Top 0.1%
9.7%
Show abstract

Background: Veterans face an elevated risk of suicide compared to the general population, motivating national efforts to develop predictive models that can guide proactive care. Current models used by the U.S. Department of Veterans Affairs (VA) rely primarily on structured electronic health record (EHR) data, though clinical notes contain rich contextual information that can be quantified using natural language processing (NLP) to derive psychosocial variables that may improve risk detection. Machine learning methods, particularly classification and regression trees (CART), can also uncover interactions between clinical and psychosocial variables, enabling identification of patient characteristics that modify suicide risk factors. However, integrating structured and unstructured data presents challenges because NLP features often greatly outnumber traditional clinical variables, potentially biasing interaction discovery. In prior work, we addressed this imbalance by introducing a weighted CART framework that balances structured variables with NLP-derived psychosocial features from semantic lexicons (SEANCE). While effective, semantic approaches summarize language into predefined constructs and may overlook important lexical variation present in clinical narratives. Methods: In this study, we extend that framework by replacing semantic features with a high-dimensional bag-of-words (BoW) representation of clinical notes and by evaluating models across cohorts defined by structured suicide risk stratification (low, medium, high) and varying temporal lookback windows. Using a cohort of 27,241 veterans, we analyzed clinical documentation collected up to 30, 90, or 270 days prior to death (or a matched index date for controls), enabling temporally flexible risk modeling. XGBoost models were trained to balance structured and unstructured features and identify cross-modal interactions between textual and clinical variables. Results: When incorporated into generalized linear models, these interactions improved predictive performance, particularly among low- and medium-risk patients, and substantially reduced the performance gap between interpretable and more complex models. Notably, the BoW representation outperformed our prior semantic index-based approach. Discussion and Conclusions: Together, these findings demonstrate the utility of interpretable NLP methods for uncovering clinically meaningful interactions between psychosocial and demographic factors in suicide risk and establish a strong benchmark for future deep learning approaches aimed at capturing richer contextual and temporal information from clinical narratives.

4
Don't stop the heart: a performance analysis of large language models and potassium dosing

Blotske, K.; Zhao, X.; Henry, K.; Murray, B.; Gao, Y.; Smith, S. E.; Wayne, N.; Ku, P.; Smith, B.; Moua, S.; Sikora, A.

2026-06-04 pharmacology and therapeutics 10.64898/2026.06.02.26354762 medRxiv
Top 0.1%
8.0%
Show abstract

Background: Electrolyte replacement is ubiquitous in the acute care setting, but its familiarity cannot belie that even small dosing errors with potassium can cause lethal cardiac arrhythmias. Recently, MedAgentBench offered a benchmark for agentic artificial intelligence (AI) including the ability to correctly dose potassium based on a single rule; however, this does not adequately reflect the clinical complexity or safety concerns of an agent that has been used as the lethal injection. The purpose of this analysis was to a probe leaderboard large language model (LLM) capabilities to follow basic dosing rules to safely replace potassium in a series of clinician-annotated cases. Methods: Using a clinician panel, we developed a series of dosing principles and 20 clinical cases reflective of the complexity of potassium replacement. External clinicians were surveyed to assess practice variability and agreement to clinician panel answers. We tested GPT-5-chat with each case in triplicate, with and without the clinician curated dosing principles, and prompted the model to answer six questions involving potassium goals, dosing, route, lab frequency, concurrent interventions, and the model's perceived level of confidence for the output and complexity of the case. The primary outcome was the rate of appropriate recommendations in comparison to clinician answers. Results: A total of 54 clinicians reviewed the 20 hypokalemia cases and hypokalemia dosing guideline. Clinicians expressed "highly agree" or "somewhat agree" for 66.8% of the cases evaluated when asked if they agree with the guideline-recommended management. When given the potassium dosing guideline, total errors dropped from 165 to 104, and average accuracy improved from 45% to 65% with GPT-5-Chat. GPT-5-Chat conveyed a high level of confidence for 100% of responses, while labeling 80% and 76% of cases as highly complex with and without the criteria, respectively. Potential harm scores were considerable in both groups, however, a notable reduction in severity scores occurred with the dosing guidance document. Recommendations on concurrent interventions and dosing had the highest rate of errors in both groups. Conclusions: Benchmarks must appropriately reflect clinical complexity to be considered valuable for the deployment of agentic artificial intelligence tools in the healthcare domain. GPT-5-Chat assessment on a comprehensive medication management task for potassium replacement showed improvement with dosing guidance, yet unfit benchmarking performance.

5
Determinants of Blood Group Antigen Expression and Prediction of Phenotypes by Machine Learning

Kranz, A.-C.; Schneider, J.; Gassner, C.; Bublitz, M.

2026-07-07 bioinformatics 10.64898/2026.07.01.735824 medRxiv
Top 0.1%
6.9%
Show abstract

Blood group antigens, defined by epitopes on the erythrocyte surface, are central to transfusion safety and maternal-fetal compatibility. While the genetic basis of many clinically relevant blood group antigens is well established, which structural and biophysical parameters determine whether a single-nucleotide variant gives rise to an antigenic phenotype remains unclear. Here, we integrate structural, biophysical, and evolutionary analyses to systematically evaluate features associated with single amino acid substitutions across 24 human protein-based blood group systems. We analyse 319 variants with curated phenotypic annotations alongside 481 control variants, identifying key determinants of null and antigenic phenotypes. Null variants are characterized by high evolutionary conservation, burial within the protein core, loss of hydrophobicity, increased polarity, and a propensity for arginine substitutions. Antigenic variants are also enriched in arginine; however, in contrast to null variants, they tend to occur at less conserved, more solvent-accessible, and structurally flexible sites. Supervised machine learning models trained on structural and biophysical descriptors were applied to distinguish (i) null and (ii) antigenic variants from controls, achieving balanced accuracies of 0.82 and 0.63, respectively. Feature importance analysis identified predicted pathogenicity, solvent accessibility, and evolutionary conservation as the most predictive determinants of null variants, whereas hydrophobicity, conservation, and flexibility dominated antigen prediction. This work establishes a framework linking molecular variation to blood group phenotypes and provides a foundation for predicting the impact of novel missense mutations in transfusion medicine and beyond.

6
Housekeeping Gene Expression Normalization in Transcriptomics Mitigates Data Leakage in Machine Learning Models

Ribas, G. T.; Riella, C. V.; Guizelini, D.; Menegatti Rigo, M.; Riella, L. V.; Borges, T. J.

2026-04-24 bioinformatics 10.64898/2026.04.24.720637 medRxiv
Top 0.1%
6.3%
Show abstract

BackgroundInappropriate normalization can lead to data leakage and overfitting in machine learning models. Accurately identifying housekeeping genes (HKGs) can overcome this problem and is crucial for normalizing gene expression data, particularly in RNA-Seq experiments. ResultsFirst, we demonstrate that the gene expression of commonly used HKGs significantly changes over time due to immunosuppressive treatments in transplant recipients. Using large public transcriptomic datasets of kidney transplantation, we developed a pipeline based on the genes coefficient of variation, stability, and Gini coefficient, and identified nine stable and better-suitable HKG candidates. Our results demonstrate that these HKGs improve the robustness and generalizability of machine learning models by minimizing data leakage, as evidenced by superior performance compared to benchmark methods like median ratio normalization and trimmed mean of M values. ConclusionsThis approach enables more accurate comparison of gene expression datasets across different clinical scenarios, improving the reliability of biomarker identification and enhancing personalized treatment strategies.

7
An AI-Powered Trisomy 21 Research Assistant

NANDI, S.; Sundararajan, Z.; Subirana-Granes, M.; Espinosa, J. M.; Pividori, M.; Sullivan, K. D.; Galbraith, M. D.; Costello, J.

2026-06-11 bioinformatics 10.64898/2026.06.08.730893 medRxiv
Top 0.1%
6.2%
Show abstract

Down syndrome, caused by trisomy 21, increases the risk of diverse co-occurring conditions. With more than 34,000 related publications indexed in PubMed as of early 2026, keeping pace with this expanding literature is challenging. While general-purpose large language models are widely used for information retrieval, they often rely on broad training data rather than specific evidence. Retrieval-augmented generation (RAG) improves rigor and reliability of responses by linking model outputs to source texts. In research, source texts are peer-reviewed articles. Standard implementations treat all manuscript sections equally, allowing background text to rank as highly as experimental results. To focus model outputs on experimentally supported responses, we developed the T21 Research Assistant, a section-aware RAG system that prioritizes Results sections to ground responses in primary experimental evidence. The system draws exclusively from 1,789 open-access Down syndrome publications from PubMed Central, including 327 NIH INCLUDE-funded studies, and uses a multistage pipeline for query validation, retrieval, reranking, synthesis, and citation verification. Built on NVIDIA Nemotron models, it generates structured, cited responses. Evaluation using expert-curated questions demonstrated strong performance, achieving a BERTScore F1 of 0.712 and recall of 0.758, comparable to or exceeding leading proprietary and open-source models. T21 Research Assistant is available at: https://bioinformatics.cuanschutz.edu/t21-res-assi/

8
Comorbidity structure as an inductive bias: Comparing output-head designs for multi-label prediction of diabetes and myocardial infarction complications

Asumboya, W. A.; Agbenorhevi, P. K.; Adams, C. F.; Ayariga, D. A.; Adjadeh, T.; Adams Ziblim, S.; Kwofie, S. K.

2026-06-23 bioinformatics 10.64898/2026.06.18.733068 medRxiv
Top 0.1%
5.7%
Show abstract

BackgroundClinical complications are often predicted with separate sigmoid outputs, even when the target labels arise from related pathophysiological processes. This paper asks whether output-layer choice should reflect both predictive convenience and the biological structure assumed among complications. The central premise is that label-dependence mechanisms are explicit hypotheses about comorbidity, not generic modelling additions. MethodsOutput-head assumptions were compared across two clinically distinct multi-label prediction tasks. In Type 2 diabetes (T2D), six heads were evaluated for nephropathy, neuropathy, and retinopathy: independent baseline, linear additive, multiplicative, symmetric conditional random field (CRF), residual multilayer perceptron (MLP), and combined additive-multiplicative. In myocardial infarction (MI), four heads were evaluated for ventricular tachycardia, ventricular fibrillation, and atrioventricular block: independent baseline, linear additive, multiplicative, and symmetric CRF. All experiments used five training data fractions and seven independent seeds, with the same shared-backbone protocol within each disease setting. ResultsIn T2D, the symmetric CRF gave the most consistent improvement pattern, ranking highest at full data and at the two lowest data fractions while adding only three interaction parameters. At 20% training data, it was the only interaction head whose aggregate mean exceeded the independent baseline. The residual MLP, despite 123 interaction parameters, remained below the baseline across all T2D fractions. In MI, rankings changed across fractions: the multiplicative head led at 80% and 60%, the CRF led at 100% and 20%, and the baseline led at 40%. The combined additive-multiplicative head did not improve robustness in T2D and showed the largest negative baseline-relative deviations at lower fractions. ConclusionThe findings support a biology-guided view of output-layer design. A small constrained mechanism was most useful when its symmetry matched the shared microvascular structure of T2D, whereas the heterogeneous electrophysiology of MI produced no stable winner. Output-layer choice should therefore be reported and defended as an assumption about disease structure instead of a routine hyperparameter decision. Author summaryMany clinical prediction models treat complications as separate outcomes, even when clinicians know they often arise together. We studied whether the last layer of a model should reflect that biological knowledge. We compared several output heads across two disease settings: Type 2 diabetes, where nephropathy, neuropathy, and retinopathy share a common microvascular origin, and myocardial infarction, where electrical complications arise from a mixture of shared and location-specific mechanisms. We found that a small symmetric CRF head was most useful in the diabetes task, especially when training data were limited, while no single interaction head dominated in myocardial infarction. This suggests that modelling comorbidity is not only a technical choice; it is a statement about how disease processes relate to one another. Our results encourage researchers to report and justify output-layer design as part of the clinical modelling argument, rather than treating it as a routine hyperparameter.

9
An electrocardiogram-based machine learning model for distinguishing complete Kawasaki disease.

Nakano, T.; Saito, K.; Noda, K.; Asai, Y.; Kojima, A.; Uchida, H.; Ohira, Y.; Ito, H.; Kawada, J.-i.; Yoshikawa, T.

2026-05-06 pediatrics 10.64898/2026.04.30.26352183 medRxiv
Top 0.1%
5.6%
Show abstract

Kawasaki disease (KD) is a systemic vasculitis in young children, and early diagnosis remains challenging when clinical features are incomplete or overlap with those of other febrile illnesses. Because electrocardiography (ECG) is noninvasive and widely available, we investigated whether ECG-derived features could help distinguish complete KD from pediatric patients with fevers. We conducted a single-center retrospective study of hospitalized febrile children aged 1-8 years who underwent digital 12-lead ECG recording during the initial evaluation. Five amplitude features and six timing features extracted from the ECG were used to develop a logistic regression model to distinguish between complete KD and other febrile illnesses. The model discriminated between the KD and non-KD groups in the validation dataset. The prediction score was not significantly correlated with the age and body temperature. S-wave amplitude, the RR interval, and P-and Q-wave amplitudes were suggested to contribute to discrimination. These findings suggest that ECG-derived features may provide adjunctive information for distinguishing complete KD from other febrile illnesses. Author SummaryKawasaki disease is an inflammatory illness in young children that can lead to coronary artery complications if treatment is delayed. Early diagnosis is often difficult because its initial symptoms overlap with those of many common febrile illnesses. We investigated whether a routine 12-lead electrocardiogram (ECG), which is noninvasive, rapid, and widely available, contains information that can help distinguish complete Kawasaki disease from other febrile conditions. We retrospectively analyzed digital ECGs from hospitalized febrile children and extracted waveform amplitude and timing features. Using these features, we built a logistic regression model and evaluated it in a temporally separate validation cohort. The model distinguished patients with Kawasaki disease from patients with fever. P-, Q-, and S-wave amplitudes and the RR interval were repeatedly selected as important contributors, suggesting that both waveform morphology and heart-rate-related information may be relevant. These findings indicate that ECG-derived features may provide useful adjunctive information during the clinical assessment of complete Kawasaki disease.

10
E-InfertilityTest: An Explainable AI Framework for Male Infertility Assessment

Das, G.; Ghosh, B.; Ghosh, Z.

2026-05-25 bioinformatics 10.64898/2026.05.21.726746 medRxiv
Top 0.1%
5.6%
Show abstract

Male infertility has emerged as a significant concern in modern society, with genetic defects as one of the major underlying cause behind it. This impairment negatively impacts sperm motility and morphology, leading to conditions such as Asthenozoospermia (reduced sperm motility), Teratozoospermia (abnormal sperm morphology) and sometimes Asthenoteratozoospermia (both motility and morphology defects). Assisted reproductive technologies (ART), such as in-vitro fertilization (IVF), offer a potential solution for such cases but with a low success rate. Classical semen analysis provides only a phenotypic snapshot without revealing the fertilizing potential of the sperms. Hence, in order to screen the functional sperm population as well as to get a deeper insight into the reasons underlying the aberrant sperm population, it is important to study their genetic profile. In this work, we have performed a meta analysis of the transcriptomic data of infertile sperms from Asthenozoospermia and Teratozoospermia patients with that from fertile sperms of normal individuals. Thereafter we have screened a signature gene set which has been used to develop a prediction model named Explainable Infertility Test (E-InfertilityTest) to classify between fertile versus infertile sperm at the preliminary level. For each prediction, it will also provide the set of genes which are playing a dominant role towards such prediction. Thus, it will provide patient specific dominant gene expression profile responsible for the aberration. This work warrants validation experiments in future to substantiate the models performance in a clinical setting. User can access the tool named E-InfertilityTest as a standalone version on GitHub. Github Linkhttps://github.com/zglabDIB/einfertility.git

11
Synthetic-data augmented calibration for expert-informed rare disease models

Yang, H.; Rachel, T.; Litwin, T.; Karakioulaki, M.; Reimer-Taschenbrecker, A.; Timmer, J.; Has, C.; Binder, H.; Hess, M.

2026-05-20 bioinformatics 10.64898/2026.05.18.725833 medRxiv
Top 0.1%
5.6%
Show abstract

Clinical data for rare diseases are sparse, noisy, and heterogeneous, complicating calibration of ordinary differential equation (ODE) models. Thus, we introduce a noise-robust calibration in latent space that combines expertderived ODEs with learned latent representations. Our approach leverages synthetic ODE trajectories, augmenting our scarce observations to train a model-specific autoencoder representation and imputer. During calibration, observed and ODE-generated trajectories are compared in latent space, and ODE parameters are updated by minimizing their latent distance. In a controlled ABCDE simulation model, the imputer outperformed a carry-forward baseline for moderate parameter shifts, parameter recovery remained stable under random missingness, calibration remained robust to additional noise variables despite reduced downstream identifiability, and distinct dynamics formed visually separable latent trajectories. On a custom developed ODE model for real Epidermolysis Bullosa patients, the calibrated phenomenological model reproduced patient-level trajectories from sparse observations. Thus, we conclude that our latent-space calibration approach supports rare-disease modeling.

12
Synthetic Data Generation and Nonparametric Techniques for Assessing Multivariate Similarity to Address Small-Sample Size Challenges

Heine, J.; Fowler, E.; Eschrich, S. A.; Schell, M.

2026-05-07 bioinformatics 10.64898/2026.05.04.722226 medRxiv
Top 0.1%
5.5%
Show abstract

Data modeling in biomedical research often operates in the small-sample regime, where the number of observations is small relative to the data dimensionality; the detrimental effects of limited sample sizes are well documented in cancer studies. Synthetic data offers a potential solution to data shortfalls provided that the data generated is an adequate facsimile of the underlying distribution; the adequacy of such synthetic data remains an open-ended problem. In this work, we evaluate a synthetic generator proposed previously. The generator applies a series of transformations to the observed data to accommodate the small-sample size resulting in an uncoupled representation, where uncorrelated marginal distributions are modeled with optimized univariate kernel density estimation. In this report, (1) we develop a nonparametric method for assessing multivariate similarity based on the Cramer-Wold theorem and random projection testing, (2) investigate when the absence of bivariate correlation approximates independence in a non-normal setting, and (3) evaluate artifacts induced by data compression. The presentation is primarily methodological; low-dimensional data were used so each stage of the generation process could be analyzed explicitly. A formal testing framework was developed by comparing random projection level outcomes with a two-sample test, modeling these outcomes as Bernoulli trials, aggregating replicate outcomes within each projection direction, and pooling outcomes across many directions, yielding a scalable standardized normal test-statistic. The key innovation was decoupling the two-sample test significance level from that governing finalized normal inference. We showed the same projection framework also evaluates the full multivariate covariance structure. The generator produced high-fidelity multivariate synthetic data when the bivariate correlation approximates independence in the non-normal setting; in highly compressed data, residual modes were best modeled as normally distributed regardless of their intrinsic distributional form. Ongoing work includes applying these methods to higher-dimensional, diverse data.

13
Glitch genes: embedding geometry predicts functional fragility in single-cell foundation models

Whalley, J. P.

2026-06-27 bioinformatics 10.64898/2026.06.22.733850 medRxiv
Top 0.1%
5.5%
Show abstract

BackgroundSingle-cell foundation models are increasingly used for perturbation prediction and gene network inference, but their learned gene representations are rarely audited directly. In natural language processing, geometric analyses of token embeddings have revealed anomalous "glitch tokens" associated with erratic model behaviour. Whether analogous representational anomalies exist in biological foundation models remains unknown. ResultsThis study introduces a weight-only geometric audit framework that scores genes by embedding norm, centroid distance, cosine similarity, and isolation to identify representational outliers. Applied to Geneformer, scGPT, and scFoundation, the analysis identifies hundreds of outliers in discrete-tokenisation models. Shared Geneformer-scGPT outliers are enriched for loss-of-function intolerance (OR=12.0) and disease association (OR=3.7), whereas scFoundations continuous value embeddings form a near-isotropic space with no detectable enrichment under the annotation panels tested. In Geneformer, geometric anomaly predicts perturbation sensitivity ({rho} = 0.725); the signal is supported by mask-in-place experiments, shows rank agreement in real PBMC cells, and correlates with Replogle perturb-seq effect sizes ({rho} = 0.645). Metric decomposition separates magnitude-driven outliers, enriched for highly expressed housekeeping genes, from isolation-driven outliers enriched for tissue-restricted genes. ConclusionsTokenisation strategy helps determine which genes are represented reliably. Embedding geometry provides a rapid, model-agnostic diagnostic that requires only an embedding matrix and can flag genes whose representations warrant caution before downstream use.

14
BioMADE: Predicting Torsades de Pointes from molecular structures through biologically informed representations

Acitores Cortina, J. M.; Schut, M. C.; Tatonetti, N. P.

2026-05-11 bioinformatics 10.64898/2026.05.06.723121 medRxiv
Top 0.1%
5.5%
Show abstract

Drug-induced arrhythmias, particularly Torsades de Pointes (TdP), pose a significant risk to patient safety and can sometimes have life-threatening outcomes. They remain a major concern in drug development and regulation. Machine learning (ML) has become a powerful tool for analyzing complex biological and chemical datasets, enabling researchers to identify subtle patterns that differentiate safe compounds from those likely to cause dangerous cardiac effects. However, most existing in silico approaches do not sufficiently incorporate biological elements, relying heavily on chemical and structural properties or on computationally expensive simulations. Here, we introduce BioMADE, a novel ML framework that harnesses small-molecule-protein activity profiles from publicly available datasets to predict TdP risk without requiring exhaustive mechanistic annotation. Activity data from ChEMBL were used to train individual models for each gene, which predict activity values for any given compound. A curated set of arrhythmia-relevant genes was then used to construct a latent biological embedding (BioMADE embedding) for each molecule. We validated the performance of these features in distinguishing biological elements such as ATC3 class, showing superior classification performance compared with representations such as Molformer (lacks biological information) and MACCS (limited chemical properties) (0.85 AUROC vs 0.81 and 0.73, respectively). BioMADE representations served as input to a support vector machine classifier to discriminate TdP-inducing drugs from safe compounds. BioMADE achieved an AUROC of 0.89 in internal validation, indicating strong predictive performance. Against state-of-the-art models such as ADMEThyst, BioMADE achieved an AUROC of 0.74 on ADMEThysts validation set (vs. 0.72 for ADMEThyst). When we combined both approaches, the AUROC reached 0.77. These results demonstrate that BioMADE provides a scalable, biology-informed, and generalizable approach for predicting drug-induced toxicities. By integrating protein activity profiles into toxicology modeling, our framework highlights the critical role of human biology in adverse drug reaction prediction, an aspect often overshadowed by purely chemical or structural descriptors.

15
A unified smoothing framework for protein domain bigram model

Cui, X.; Iyer, G.; Durand, D.

2026-06-18 bioinformatics 10.64898/2026.06.14.732219 medRxiv
Top 0.1%
5.5%
Show abstract

MotivationBiomolecular sequences can be represented as strings over an alphabet, an analogy that has motivated many applications of computational linguistic techniques to biological problems. However, such methods must be adapted to the characteristic scale and organization of biomolecular data. Here, we consider the problem of bigram smoothing for multidomain protein architectures, where domain bigram frequency data is extremely sparse and differs from textual data in alphabet size, string length distribution, the relationship between bigram and unigram frequencies, tandem repeat lengths, and the distribution of domain adjacencies. Moreover, some domain combinations are unobserved because they are biologically incompatible, others because the data are incomplete. A smoothing method that distinguishes these two cases is required. ResultsWe propose a unified smoothing framework based on interpolation that can be tuned to accommodate different bigram data characteristics. Within this framework, we design specific model variants suited to protein domain bigram data: these assign low adjusted counts to pairs that are likely incompatible, while making appropriate adjustments for undersampled pairs. We demonstrate empirically that this approach distinguishes the two cases while preserving the characteristic signatures of multidomain data. Availability and implementationImplementations of smoothing methods, the scripts used to generate all results presented in this paper, and the curated lists of extracellular and DNA-binding domains are available at https://codeberg.org/xcui297/protein-domain-smoothing.

16
Evaluation of AI-Generated Synthetic Data for Clinical Research in Secondary Cardiovascular Prevention among Dyslipidemia Patients

Bonomi, A.; Werba, J. P.; Saccani, S.; Lu, L. L.; Coser, A.; Franchi, M.; Valsecchi, C.; Teruzzi, E.; Terragni, A.; Centenaro, C.; Scatigna, M.; Pompilio, G.

2026-06-15 cardiovascular medicine 10.64898/2026.06.12.26355456 medRxiv
Top 0.1%
5.5%
Show abstract

Background: Access to high-quality clinical data is essential for advancing medical research and developing effective medical statistical and Artificial Intelligence models. However, privacy regulations and logistical barriers often hinder timely access to real-world data. Synthetic data offer a promising solution, preserving the statistical characteristics of original datasets while protecting patient privacy. Objectives: This study investigates the use of synthetic data for secondary cardiovascular prevention in patients with dyslipidemia, using two real-world datasets from Centro Cardiologico Monzino. Methods: Given the high dimensionality and limited sample size of the datasets, we employed a custom generative framework based on Large Language Models (LLMs). Pre-trained LLMs were fine-tuned on original clinical records to synthesize tabular data replicating source-data distributions. Fine-tuning was performed within the Centro Cardiologico Monzino's secure infrastructure to ensure data sovereignty. We evaluate clinical utility and privacy using fidelity and privacy metrics, identifying the optimal generative model and benchmarking against traditional anonymization methods. Results: Synthetic data achieved a superior trade-off than classically anonymized datasets. Real and synthetic datasets showed strong agreement, with significant distributional differences limited to few variables. Models trained on synthetic data replicated key associations from the original dataset, including therapy modification and creatine phosphokinase as predictors of SAMS, and pharmacological intensity as the main driver of LDL-C reduction. Conclusions: Results support the feasibility of using synthetic data as a proxy for real-world datasets in exploratory analyses and model development. Despite slight attenuation of some effect sizes, preserved clinical relationships reinforce the validity of synthetic data in medical research.

17
Quantifying Evidence for Competing Biomedical Hypotheses using Large Language Models and Bayesian Analysis

Moore, B. M.; Freeman, J.; Millikin, R. J.; Mohanty, C.; George, K. S.; Bal, A.; Lock, C.; Sauer, J.-D.; Spurgeon, M. E.; Moore, D. L.; Travers, B. G.; Stewart, R.

2026-06-07 bioinformatics 10.64898/2026.06.05.730173 medRxiv
Top 0.1%
5.5%
Show abstract

Science fundamentally depends on the generation and testing of hypotheses, many of them controversial. An explosion in scientific literature has made evaluating hypotheses even within a domain a problem of scale, and risks slowing an already extensive consensus-building process. While this challenge has prompted interest in automated hypothesis evaluation tools, existing methods have not yet proven effective for comparing hypotheses. Here, we introduce KM-GPT-DCH, an algorithm that combines co-occurrence methods with large language models (LLMs) to develop a transparent and reproducible literature-based algorithm to compare controversial hypotheses using a structured scoring approach with Bayesian methods to estimate confidence. When testing the algorithm on historical controversial hypotheses previously decided, KM-GPT-DCH chooses the correct hypothesis with high confidence several years before the scientific community or public do so. We further apply the algorithm to compare twenty unresolved controversial hypothesis pairs providing guidance for future research. The method can help researchers and the public to evaluate biomedical hypotheses such as "Is it more likely that monoamine deficiency or inflammation causes depression?" It can also be used to assess and visualize historical trends in the scientific literature. A web-based implementation of the algorithm is freely available at https://skim.morgridge.org.

18
Corpus-wide causality: Algorithm design & application for aggregating gene-disease causal evidence

Bansal, N.; Parsodkar, A. P.; Pathak, A.; Narayanan, M.

2026-05-12 bioinformatics 10.64898/2026.05.08.723796 medRxiv
Top 0.1%
5.4%
Show abstract

Identifying causal relationships and distinguishing them from associations is a central scientific endeavor with many applications; knowing causal links between genes and diseases, for instance, can focus drug discovery on curing diseases beyond just symptom management. Despite several studies on automatically extracting relations between entities from large biomedical literature corpora like PubMed, only a few studies extract causal relations from abstracts and even fewer summarize corpus-level evidence for causal links. Recently, Large Language Models (LLMs) have been increasingly deployed to summarize biomedical information and extract relations; however, there is a distinct lack of explicit benchmarking comparing these generalized LLM-based methods against specialized, domain-aware frameworks for corpus-wide causal inference. In this work, we develop a method to infer Corpus-Wide Causal Score (CWCS) of a gene-disease (G-D) pair by integrating two pieces of evidence: (i) network-based causal signals in a prior gene regulatory network, quantified as a CWCS-Net score using an existing multilayer network centrality algorithm; and (ii) corpus-wide literature evidence, quantified as a CWCS-TD (TD for Truth Discovery) score using a newly-developed TD algorithm. Our CWCS-TD (scoring) algorithm jointly and iteratively estimates causal scores for multiple G-D pairs while modeling the reliability of PubMed abstracts co-mentioning them; and represents an advance in the field of TD algorithms due to its incorporation of bibliometric features of publications to address the challenge of sparsity of abstracts that assert a G-D causal relation. Using OMIM as an external expert-curated reference to evaluate classifications of G-D pairs as causal or not, our CWCS method achieved a causal class F1 score of 0.600 across ten diseases, outperforming both LLMs, GPT-4o and MMed-Llama 3 (this performance trend also persists when using area under the precision-recall curve as the evaluation metric). Both LLMs exhibit high recall accompanied by comparatively low precision, resulting in lower causal class F1 scores (0.505 for GPT-4o and 0.522 for MMed-Llama 3) due to large number of false positive predictions. Taken together, these evaluations and other ablation studies show the promise of our carefully designed algorithm in collating and integrating evidence of biomedical causal relations from both network- and literature-based sources, thereby supporting its broader applicability.

19
Widespread use of invalid statistical tests in biomedical machine learning

Zeng, T.; Li, H.; Zhang, S.; Tan, Y. Q.; Tian, F.; Orban, C.; An, L.; Che, W.; Cheng, J.; Chong, J. S. X.; Dehestani, N.; Dong, Z.; Li, X.; Li, Z.; Lim, M. J. R.; Lin, Y.; Ling, Q.; Ling, Z.; Low, X. Z.; Mansour L., S.; Ng, K. K.; Nguyen, T. T.; Ooi, L. Q. R.; Pande, S.; Qian, X.; Ruan, J.; Wang, Z.; Xie, Y.; Zhang, C.; Zhang, Y.; Patil, K.; Parkes, L.; Dhamala, E.; Chopra, S.; Zalesky, A.; Holmes, A.; Eickhoff, S.; Zhou, J. H.; Renaud, O.; Dosenbach, N.; Kording, K. P.; Bzdok, D.; Nichols, T.; Yeo, B. T. T.

2026-05-20 bioinformatics 10.64898/2026.05.17.724301 medRxiv
Top 0.1%
5.4%
Show abstract

Machine learning is accelerating biomedical research. Cross-validation is widely used to compare predictive performance - not only to benchmark algorithms, but also to inform scientific applications, such as ranking biomarkers. However, prediction performance estimates across cross-validation folds are not independent. Standard tests for comparing prediction performance (e.g., paired t-test) assume independence and can therefore inflate false positive rates. In a PRISMA-guided meta-analysis of 210 studies (impact factor [≥]15, 1 June 2020 - 1 June 2025), we find that 97% ignored fold dependence when comparing prediction performance. This problem is ubiquitous across scientific fields and unaffected by impact factor, rigor-promoting policies, or open science practices. Simulations across 420 scenarios spanning four diverse datasets show that ignoring fold dependence leads to invalid false positive control in most settings. Repeated cross-validation further compounds this problem, with false positive rates rising toward 100% as the number of repetitions grows. Existing fold-dependence-aware tests rely on strong assumptions because the variance of fold-level statistics and the between-fold correlation cannot be disentangled under standard cross-validation. We therefore propose the SHARP (Split-HAlf RePeated) test, a simple modification to standard cross-validation that enables direct estimation of variance and correlation. Benchmarked against 12 tests, SHARP provides the best overall balance of false-positive control, statistical power, and confidence-interval calibration across simulation schemes. We conclude by providing best practices and reporting guidelines for valid model comparison inference in biomedical machine learning and beyond.

20
CausalKnowledgeTrace: A Novel Computational Framework for Automated Literature-Based Causal Graph Construction and Evidence-Based Variable Selection in Biomedical Research

Upadhayaya, R.; Pradhan, M. M.; Metzger, V. T.; Malec, S. A.

2026-05-12 bioinformatics 10.64898/2026.05.07.723601 medRxiv
Top 0.1%
5.2%
Show abstract

BackgroundVariable selection for causal inference from observational biomedical data is challenging, as overlooking confounders or conditioning on colliders leads to biased estimates. While vast causal knowledge exists in biomedical literature, manually extracting this information for principled variable selection is impractical at scale. MethodsWe developed CausalKnowledgeTrace, a Python-based computational framework with Django web interface that systematically leverages structured causal knowledge from the Semantic MEDLINE Database (SemMedDB) to inform variable selection in causal studies. The system implements a six-stage analysis pipeline using NetworkX for graph operations, including graph parsing, basic analysis, comprehensive cycle detection, systematic generic node removal, post-removal analysis, and formal causal inference with bias detection. ResultsAnalysis of the hypertension-Alzheimers relationship across three degree neighborhoods (1-3) demonstrated systematic scaling of causal complexity: 361-866 variables, 429-1,442 relationships, with graph densities of 0.0033-0.0019. The analysis revealed complex cyclic structures with 54-606 baseline cycles across degree levels. Processing times ranged from 0.3-1.0 seconds for all three degrees, demonstrating computational efficiency for complex biomedical networks. Key confounders identified across all degrees included inflammation, diabetes, insulin resistance, obesity, and ischemia. In the third degree of graph, the pipeline structurally identified 39 confounders, 11 mediators, and 3 colliders from the causal graph. Among the key identified confounders and mediators--including obesity, oxidative stress, ischemia, and vascular diseases--all were found to have strong supporting evidence in established epidemiological and pathophysiological literature. ConclusionsCausalKnowledgeTrace provides a scalable, evidence-based approach to causal graph construction that systematically identifies confounders and bias structures often missed by conventional approaches. The Python-Django architecture enables both standalone analysis and integration into larger computational workflows, representing a significant advance in computational support for causal inference in biomedical research. Statement of SignificanceO_ST_ABSProblem or IssueC_ST_ABSSelecting proper confounders and variables for causal inference from observational biomedical datasets is challenging and often biased by limited expertise or manual review. What is Already KnownExisting approaches rely on domain experts, statistical variable screening, or manual construction of causal graphs, but these often overlook literature-documented confounders and complex biases. What this Paper AddsThis paper introduces an automated, literature-based framework for synthesizing and validating causal graphs, identifying critical variables and complex bias structures, such as M-bias and butterfly bias, with full evidentiary traceability. Who would benefit from the new knowledge in this paper?Epidemiologists, biomedical researchers, informaticians, and clinical investigators seeking reliable and transparent causal modeling for observational studies.